Papers with argument quality
Mining, Assessing, and Improving Arguments in NLP and the Social Sciences (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | a tutorial on computational argumentation is updated to address the problem of argument quality . argument quality is a field of interdisciplinary research that connects natural language processing to social sciences . |
| Approach: | They present an updated version of the EACL 2023 tutorial on argument quality . they will focus on the notions of argument quality across disciplines . |
| Outcome: | The updated version of the EACL 2023 tutorial focuses on argument quality assessment . the authors will focus on the interface between Argument Mining and Deliberation Theory . |
Towards Argument Mining for Social Good: A Survey (2021.acl-long)
Copied to clipboard
| Challenge: | Argument Mining is a social science-based approach to analysis and analysis of arguments. |
| Approach: | They propose a novel definition of argument quality which integrates the social science literature and the argument quality. |
| Outcome: | The proposed definition of argument quality integrates the social science literature and the argument quality debate. |
Bridging Argument Quality and Deliberative Quality Annotations with Adapters (2023.findings-eacl)
Copied to clipboard
| Challenge: | Assessing the quality of an argument is a complex, highly subjective task . argument quality dimensions are complex and dependent on the context in which it is assessed . |
| Approach: | They propose a multi-task learning framework that incorporates knowledge about related dimensions into the learning process. |
| Outcome: | The proposed framework improves quality prediction in an extrinsic, out-of-domain task. |
ConQRet: A New Benchmark for Fine-Grained Automatic Evaluation of Retrieval Augmented Computational Argumentation (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for evaluating RAArg are costly and lack long, complex arguments and real-world evidence. |
| Approach: | They propose to use multiple fine-grained LLM judges to evaluate RAArg using a new benchmark that features long and complex human-authored arguments on debated topics. |
| Outcome: | The proposed methods provide better and more interpretable assessments than traditional single-score metrics and even previously reported human crowdsourcing. |
Towards a Perspectivist Turn in Argument Quality Assessment (2025.naacl-long)
Copied to clipboard
| Challenge: | Argument quality is a key aspect of computational argumentation (CA), but it still exhibits a high degree of subjectivity in perception. |
| Approach: | They propose to use a multi-layered classification to target two aspects of argument quality in a systematic review of NLP datasets. |
| Outcome: | The proposed model improves the quality of annotators and their ability to be used in perspectivist research. |
Efficient Pairwise Annotation of Argument Quality (2020.acl-main)
Copied to clipboard
| Challenge: | Especially crowdsourcing suffers from assessors having different reference frames to base their judgments on and task instructions being nondescript and therefore unhelpful in ensuring consistency. |
| Approach: | They propose an efficient annotation framework for argument quality that uses a stochastic transitivity model and an effective sampling strategy to infer high-quality labels. |
| Outcome: | The proposed model significantly outperforms existing annotation procedures and offers statistical insights into argument quality. |
A Multi-persona Framework for Argument Quality Assessment (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for argument quality assessment do not consider multi-perspective evaluation due to subjective nature of arguments. |
| Approach: | They propose a multi-persona framework for argument quality assessment that simulates diverse evaluator perspectives through large language models. |
| Outcome: | The proposed framework outperforms baselines while providing comprehensive multi-perspective rationales on IBM-Rank-30k and IBM-ArgQ-5.3kArgs datasets. |
Argument Mining for Review Helpfulness Prediction (2022.emnlp-main)
Copied to clipboard
| Challenge: | Argumentational features have been shown to be promising indicators of product review helpfulness, but their utility has been limited due to the lack of resources and large-scale experiments investigating their utility. |
| Approach: | They present an argumentational argumentation model that annotates 878 Amazon reviews on headphones and uses it to evaluate argument quality. |
| Outcome: | The proposed model improves the state-of-the-art model under text-only and text-and-image settings. |
Contextual Interaction for Argument Post Quality Assessment (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for assessing the quality of natural language arguments are limited . existing methods focus on evaluating individual argument posts, but they often fail to distinguish between arguments with a narrow quality gap. |
| Approach: | They propose to use supervised contrastive learning to model arguments' quality . large language models with in-context examples harness the power of LLMs . |
| Outcome: | The proposed approach outperforms state-of-the-art models on a publicly available dataset . it shows that the LLMs with in-context examples are more effective than baseline models . |
FORECAST2023: A Forecast and Reasoning Corpus of Argumentation Structures (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing work on the role of reasoning in forecasting has focused on surface-level features such as linguistic markers, the use of comparison classes, and overall dialectical complexity. |
| Approach: | They propose to use a dataset of such prediction rationales to create a fully automated annotation system that can be used to enhance the argumentation. |
| Outcome: | The proposed dataset provides a uniquely fine-grained and close characterisation of the structure of argumentation with potential impact on forecasting domains from intelligence analysis to investment decision-making. |
Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It’s Best to Relate Perspectives! (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to subjectivity in natural language processing are subjective . authors argue that disagreement should not be regarded as a problem . |
| Approach: | They propose to account for subjective perspectives of individuals and objective concepts that build a common ground between annotators. |
| Outcome: | The proposed architectures increase the averaged annotator-individual F1-scores up to 43% over a majority-label model. |
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks (2026.findings-acl)
Copied to clipboard
| Challenge: | Argumentation skills are an essential toolkit for large language models (LLMs). |
| Approach: | They propose a benchmark to evaluate the generalizability of five LLM families across 46 computational argumentation tasks. |
| Outcome: | The proposed benchmark evaluates the generalizability of five LLM families across 46 computational argumentation tasks covering mining arguments, assessing perspectives, evaluating argument quality, reasoning about arguments, and generating arguments. |